Review



optimized hardware streaming architecture for the cnn  (Xilinx Inc)

 
  • Logo
  • About
  • News
  • Press Release
  • Team
  • Advisors
  • Partners
  • Contact
  • Bioz Stars
  • Bioz vStars
  • 90

    Structured Review

    Xilinx Inc optimized hardware streaming architecture for the cnn
    The architecture of <t>Convolutional</t> <t>Neural</t> <t>Network</t> <t>(CNN)-based</t> hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.
    Optimized Hardware Streaming Architecture For The Cnn, supplied by Xilinx Inc, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
    https://www.bioz.com/product/optimized+hardware+streaming+architecture+for+the+cnn/optimized+hardware+streaming+architecture+for+the+cnn/pmc07288095-410-10-16
    Average 90 stars, based on 1 article reviews
    optimized hardware streaming architecture for the cnn - by Bioz Stars, 2026-09
    90/100 stars

    Images

    1) Product Images from "Real-Time Energy Efficient Hand Pose Estimation: A Case Study"

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study

    Journal: Sensors (Basel, Switzerland)

    doi: 10.3390/s20102828

    The architecture of Convolutional Neural Network (CNN)-based hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.
    Figure Legend Snippet: The architecture of Convolutional Neural Network (CNN)-based hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.

    Techniques Used: Activation Assay

    Design process overview; the first box illustrates the software-level design phase, while the other boxes illustrate the hardware related design phase. The first stage is quantization-aware training (QAT) in which we decrease the CNN memory demand as well as the computation time on the hardware. The second stage is hardware streaming architecture (HSA) where the underlying hardware structure is designed for the CNN. In system integration (SI) stage, the programmable logic PL and the processing system PS are brought together and the interface with the memory is configured through the hardware system integration sub-stage. Furthermore, the on-Chip Software is developed for preprocessing and interfacing with the peripherals.
    Figure Legend Snippet: Design process overview; the first box illustrates the software-level design phase, while the other boxes illustrate the hardware related design phase. The first stage is quantization-aware training (QAT) in which we decrease the CNN memory demand as well as the computation time on the hardware. The second stage is hardware streaming architecture (HSA) where the underlying hardware structure is designed for the CNN. In system integration (SI) stage, the programmable logic PL and the processing system PS are brought together and the interface with the memory is configured through the hardware system integration sub-stage. Furthermore, the on-Chip Software is developed for preprocessing and interfacing with the peripherals.

    Techniques Used: Software

    Streaming Architecture. Each CNN layer is mapped into a hardware block, and the hardware blocks are connected to each others via stream channels. The bitwidth of each stream is shown on this figure.
    Figure Legend Snippet: Streaming Architecture. Each CNN layer is mapped into a hardware block, and the hardware blocks are connected to each others via stream channels. The bitwidth of each stream is shown on this figure.

    Techniques Used: Blocking Assay

    Hardware system integration; AXI-Lite interface provides the interconnection between the PS and the PL. DMA module is integrated in the PL. This module is responsible for converting the memory mapped input to AXI stream CNN input, as well as converting the AXI stream CNN output to a memory mapped output.
    Figure Legend Snippet: Hardware system integration; AXI-Lite interface provides the interconnection between the PS and the PL. DMA module is integrated in the PL. This module is responsible for converting the memory mapped input to AXI stream CNN input, as well as converting the AXI stream CNN output to a memory mapped output.

    Techniques Used:

    Related Articles

    Activation Assay:

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study
    Article Snippet: and effectively quantizing the network model parameters, which resulted in a significantly compressed model for a negligible decrease in accuracy. .. Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA. .. Our results have shown that the FPGA achieves better performance than its software counterpart, which run on high-performance GPU for the same batch s

    Software:

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study
    Article Snippet: and effectively quantizing the network model parameters, which resulted in a significantly compressed model for a negligible decrease in accuracy. .. Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA. .. Our results have shown that the FPGA achieves better performance than its software counterpart, which run on high-performance GPU for the same batch s

    Blocking Assay:

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study
    Article Snippet: and effectively quantizing the network model parameters, which resulted in a significantly compressed model for a negligible decrease in accuracy. .. Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA. .. Our results have shown that the FPGA achieves better performance than its software counterpart, which run on high-performance GPU for the same batch s



    Similar Products

    90
    Xilinx Inc optimized hardware streaming architecture for the cnn
    The architecture of <t>Convolutional</t> <t>Neural</t> <t>Network</t> <t>(CNN)-based</t> hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.
    Optimized Hardware Streaming Architecture For The Cnn, supplied by Xilinx Inc, used in various techniques. Bioz Stars score: 90/100, based on 1 PubMed citations. ZERO BIAS - scores, article reviews, protocol conditions and more
    https://www.bioz.com/product/optimized+hardware+streaming+architecture+for+the+cnn/optimized+hardware+streaming+architecture+for+the+cnn/pmc07288095-410-10-16
    Average 90 stars, based on 1 article reviews
    optimized hardware streaming architecture for the cnn - by Bioz Stars, 2026-09
    90/100 stars
      Buy from Supplier

    Image Search Results


    The architecture of Convolutional Neural Network (CNN)-based hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.

    Journal: Sensors (Basel, Switzerland)

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study

    doi: 10.3390/s20102828

    Figure Lengend Snippet: The architecture of Convolutional Neural Network (CNN)-based hand pose estimation algorithm. The CNN takes a 128 × 128 input preprocessed image. It consists of 3 convolution layers each followed by an ReLU activation and max pooling. Afterwards, there are 2 fully connected layers with ReLU activation. A third fully connected layer (joint regression layer) regresses 3D joint positions. Conv stands for a convolution layer. Each conv is followed by a ReLU activation. FC denotes a fully connected layer.

    Article Snippet: Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA.

    Techniques: Activation Assay

    Design process overview; the first box illustrates the software-level design phase, while the other boxes illustrate the hardware related design phase. The first stage is quantization-aware training (QAT) in which we decrease the CNN memory demand as well as the computation time on the hardware. The second stage is hardware streaming architecture (HSA) where the underlying hardware structure is designed for the CNN. In system integration (SI) stage, the programmable logic PL and the processing system PS are brought together and the interface with the memory is configured through the hardware system integration sub-stage. Furthermore, the on-Chip Software is developed for preprocessing and interfacing with the peripherals.

    Journal: Sensors (Basel, Switzerland)

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study

    doi: 10.3390/s20102828

    Figure Lengend Snippet: Design process overview; the first box illustrates the software-level design phase, while the other boxes illustrate the hardware related design phase. The first stage is quantization-aware training (QAT) in which we decrease the CNN memory demand as well as the computation time on the hardware. The second stage is hardware streaming architecture (HSA) where the underlying hardware structure is designed for the CNN. In system integration (SI) stage, the programmable logic PL and the processing system PS are brought together and the interface with the memory is configured through the hardware system integration sub-stage. Furthermore, the on-Chip Software is developed for preprocessing and interfacing with the peripherals.

    Article Snippet: Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA.

    Techniques: Software

    Streaming Architecture. Each CNN layer is mapped into a hardware block, and the hardware blocks are connected to each others via stream channels. The bitwidth of each stream is shown on this figure.

    Journal: Sensors (Basel, Switzerland)

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study

    doi: 10.3390/s20102828

    Figure Lengend Snippet: Streaming Architecture. Each CNN layer is mapped into a hardware block, and the hardware blocks are connected to each others via stream channels. The bitwidth of each stream is shown on this figure.

    Article Snippet: Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA.

    Techniques: Blocking Assay

    Hardware system integration; AXI-Lite interface provides the interconnection between the PS and the PL. DMA module is integrated in the PL. This module is responsible for converting the memory mapped input to AXI stream CNN input, as well as converting the AXI stream CNN output to a memory mapped output.

    Journal: Sensors (Basel, Switzerland)

    Article Title: Real-Time Energy Efficient Hand Pose Estimation: A Case Study

    doi: 10.3390/s20102828

    Figure Lengend Snippet: Hardware system integration; AXI-Lite interface provides the interconnection between the PS and the PL. DMA module is integrated in the PL. This module is responsible for converting the memory mapped input to AXI stream CNN input, as well as converting the AXI stream CNN output to a memory mapped output.

    Article Snippet: Afterwards, we provided an optimized hardware streaming architecture for the CNN which was then implemented on Xilinx UltraScale+ MPSoC FPGA.

    Techniques: